Biomedical Physics & Engineering Express
○ IOP Publishing
Preprints posted in the last 30 days, ranked by how well they match Biomedical Physics & Engineering Express's content profile, based on 11 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.
Smid, J.; Jezdik, P.; Kalina, A.; Kudr, M.; Janca, R.
Show abstract
Background: Precise localisation of intracranial electrode contacts is essential for the interpretation of stereoelectroencephalography recordings and planning epilepsy surgery. In current clinical practice, this is typically a manual process, which is time-consuming and prone to variability. Existing automated solutions are often fragmented across multiple tools requiring technical expertise, limiting their adoption in routine clinical workflows. This study presents an open-source extension for 3D Slicer that provides an integrated, user-friendly standalone solution for the direct automatic detection of electrode contacts within a widely used medical imaging platform. Results: The proposed method combines anchor bolt-based initialisation, probabilistic segmentation of electrode structures, and non-linear modelling to precisely track true electrode trajectories. The approach was evaluated on a dataset comprising 78 cases from 73 patients, including 1,078 electrodes with 14,480 contacts. The method achieved high localisation accuracy, with a median (interquartile range) deviation of 0.10 (0.06, 0.15) mm. Only 7/1078 (0.65%) electrodes required manual correction; these specific cases were handled using tools provided within the proposed extension. Conclusions: The presented extension enables fast, accurate, and reproducible electrode contact localisation within a single integrated environment. By combining automation with intuitive user interaction, it significantly reduces processing time while maintaining clinical reliability. The tool's free availability as an extension in 3D Slicer lowers the barrier to adoption and supports the standardisation of workflows across clinical and research centres.
Zareian, B.; Fontaine, K.; Bini, J.
Show abstract
Background. Roughly, half of new type 1 diabetes (T1D) diagnoses occur in individuals under 18 years old and represent a more aggressive destruction of beta cell mass (BCM). [11C]-(+)-PHNO positron emission tomography (PET) imaging is used to assess BCM, but current pancreas PET imaging protocols are limited to adults. Previously published full count data from six healthy controls and five T1Ds (6M/5F; 22 to 53 years old) were used for retrospective analysis. Dynamic [11C]-(+)-PHNO PET/CT scans were acquired and reconstructed using full-count list-mode data. For the current comparison to full count data, 50%, 25% and 10% down-sampled count data were re-reconstructed. Pancreas and spleen (reference region) time-activity-curves (TACs) were assessed, and volume of distribution (VT, mL/cm3) was estimated using the reversible 1-tissue compartment model (1TC) with tmax of 30 min for all count levels. Pancreas and Spleen VT estimates (1TC; tmax= 30 min) were used to calculate non-displaceable binding potential (BPND) and were then correlated to semi-quantitative methods of standardized uptake value ratio (SUVR-1) (20-30 min; ref: spleen) to examine simplified methods using simulated low dose protocols. Finally, we performed dosimetry in adult, adolescent and pediatric phantoms to assess radiation dose for simulated low-dose protocols. Results. Qualitatively, increasing noise can be visualized at successive reduced-count levels images, compared to full-count images. Despite progressively increasing noise in reduced-count images, TACs at each reduced-count level remained similar to full-count TACs in both HC and individuals with T1D. Quantitatively, 1TC VT estimates were similar for all reduced count levels and range of tmax values, compared to full-count (all R2[≥]0.99). Pancreas SUVR-1 (20-30 min) and pancreas BPND (tmax = 30; ref: spleen) were highly correlated for all count levels (all R2[≥]0.80). All age groups were under both the yearly occupational and research scan radiation dose limits when examining mean effective dose equivalent with reduced (1/10th) injected dose protocols. Conclusion. Low-count reconstructed data and simplified reference region approaches provide accurate quantification compared to full-count reconstructions. These results provide evidence that it is possible to perform accurate quantification using simulated low dose protocols to quantify BCM for use in individuals with T1D under 18 years old.
Oyarzun Silva, R.; Hernandez Hernandez, P.
Show abstract
Background. Accurate delineation of the gross tumour volume (GTV) - primary tumour (GTVp) and nodal disease (GTVn) - on FDG-PET/CT is a critical step of head and neck radiotherapy planning. Comparisons between lightweight custom networks and the auto-configured nnU-Net v2 are usually reported as end-to-end pipelines, conflating the contribution of the network with that of the inference-time post-processing applied on top of it. We separated the two. Methods. MiniUNet3D (custom 3D U-Net, 18.3 M parameters) and nnU-Net v2 (3d_fullres, 88.2 M parameters) were trained on the same 578 FDG-PET/CT cases (85/15 author-defined split of the HECKTOR 2025 Task 1 set, 8 centres) and evaluated on the same internal cohort. Three arms were compared pairwise: MiniUNet3D raw output at a fixed 0.5 threshold, MiniUNet3D with a locked adaptive post-processing pipeline, and nnU-Net v2. Comparisons used paired Wilcoxon tests with bootstrap confidence intervals, Bonferroni and Benjamini-Hochberg correction, and Cohen's d; catastrophic failure (Dice < 0.01) was compared with an exact McNemar test. Cases with an empty reference for a given target were excluded from that target's analysis (n = 98 GTVp, n = 93 GTVn). Results. With post-processing matched off, nnU-Net v2 was superior: median GTVp Dice 0.799 versus 0.592 (mean difference -0.244, 95 % CI -0.300 to -0.191; d = -0.88) and GTVn 0.774 versus 0.598 (d = -0.82). Post-processing raised MiniUNet3D to 0.800 (GTVp) and 0.738 (GTVn), recovering 79 % of that difference. Post-processed, MiniUNet3D matched nnU-Net v2 on GTVp Dice (p = 0.113) but remained inferior on nodal disease after Bonferroni correction (Dice p = 0.041; surface Dice p = 0.049). Catastrophic GTVp failures were 25/98 raw, 8/98 post-processed and 1/98 for nnU-Net v2 (McNemar p = 0.016). Inference took 34 s versus 78 s per case on the same GPU. Conclusions. Post-processing recovered most, but not all, of the difference between the two models, and it did not confer robustness: an eight-fold higher rate of empty contours on small primaries persisted, which is the more consequential difference for planning safety. Pipeline comparisons reported without a post-processing ablation risk attributing to a network what post-processing supplied.
Stöhrmann, P.; Ponce de Leon, M.; Dörl, G.; Milz, C.; Graf, S.; Eggerstorfer, B.; Murgas, M.; Reed, M. B.; Falb, P. C.; Al Barede, K.; Nics, L.; Rasul, S.; Hacker, M.; Lanzenberger, R.; Hahn, A.
Show abstract
Purpose: Attenuation correction (AC) of PET images is essential for accurate quantification. Brain PET studies comprising simultaneous EEG (PETEEG) may suffer from metal artifacts in CT images (CTEEG), or improper correction when electrodes are not present in the CT (CT0). As these influences are not well-characterized, we aim to compare metal artifact reduction (MAR) techniques for CTEEG images, and evaluate differences between attenuated-corrected PETEEG using CT0 and CTEEG with MAR, synthetically placed electrodes (CTEEG-synth) and extended Hounsfield unit (HU) range. Methods: 19 healthy participants underwent two total-body PET/CT scans with [18F]FDG, with and without 32 EEG scalp electrodes, respectively. We evaluated five MARs to reduce streaks caused by the EEG electrodes in the CTEEG. Finally, CT0, CTEEG with (CTEEG-iMAR-Ext) and without extended HU range (CTEEG-iMAR) and CTEEG-synth were used to perform attenuation correction of PETEEG. We compared our results to PET0/CT0 scan using relative differences. Results: CTEEG and CTEEG-iMAR showed the smallest differences to CT0. PETEEG/CTEEG-iMAR-Ext exhibited the lowest differences to PET0/CT0 (average bias across all regions of -0.46%), followed by similar performance of PETEEG/CTEEG-iMAR (-0.73%) and PETEEG/CTEEG (-0.76%). Conversely, PETEEG/CT0 demonstrated the largest average differences (-1.81%), with values reaching -2.71% in the parietal lobe. These differences were consistent across subjects, yielding significant effects in most of the brain (pFWE < 0.05). CTEEG-synth performed not as good as CTEEG (-1.21%). Conclusions: CTEEG with extended HU range is most suitable for attenuation correction of PETEEG images, with MAR correction offering little additional improvement.
Hsiao, N.; Clifford, M.; Lin, S.-Z.; Premasiri, S.; Roots, J.; Allen, H.; Robertson, A. P.; Moafa, K.; Wardle, J.; Edwards, C.
Show abstract
Objective To evaluate the effect of vendor-integrated AI-assisted abdominal ultrasound software on operational efficiency and sonographer workload compared with manual scanning. Methods In this prospective randomised crossover study (January to February 2026), 32 healthy adults each underwent two upper abdominal examinations, one manual and one using vendor-integrated AI software (AI Abdomen Release 3.5; ACUSON Sequoia), in randomised order by two experienced sonographers, each participant scanned once by each sonographer. Scan time, hand-console interaction (keystrokes, hand travel, hover, jerk) from a custom depth-camera hand-tracking system, and operator modifications to AI outputs were recorded. Workload was assessed after each scan with the weighted NASA Task Load Index (NASA-TLX). Analysis used linear mixed-effects models. Results AI-assisted scanning reduced scan time (52.4 s, approximately 9%; 95% CI 23.7 to 81.2; P = 0.001), keystrokes (55, approximately 28%; P < 0.001) and hand travel (4.57 m, approximately 39%; P < 0.001), although the time saving was concentrated in one sonographer. Weighted NASA-TLX did not differ between conditions (-3.9 points; 95% CI - 9.3 to 1.5; P = 0.17), but subscale analyses showed reductions in mental demand (- 6.3; P = 0.03) and effort (- 7.0; P = 0.04), with no compensating increases. Sonographers modified 48 of 184 AI-generated values. Conclusion AI assistance improved operational efficiency and reduced self-reported mental demand and effort, with no compensating increase on other subscales. Gains arose under a controlled, abbreviated protocol in healthy volunteers and varied between operators, and are better read as a reshaping of operator work than its removal.
Tang, D.; Swenson, C.; Small-Zlochower, S.; Bizik, G.; Christensen, L. M.; Knösche, T.; Haueisen, J.; Ludwig, R.; Nunez Ponasso, G. C.; Noetscher, G.; Deng, Z.-D.; Makaroff, S. N.
Show abstract
Objective: Low-intensity transcranial magnetic stimulation (LI-TMS) is being investigated as a gel-free alternative to transcranial electrical stimulation (tES), but existing systems remain almost exclusively single-channel and cannot electronically steer the induced electric field. We present the design, modeling, and experimental measurement of a wearable whole-head, multichannel, steerable LI-TMS array. Methods: The system comprises a 102-channel conformal coil array with independently controlled drivers capable of arbitrary waveform synthesis, together with a boundary element fast multipole method (BEM-FMM) framework that computes the coil currents required to produce prescribed cortical field patterns. A 12-channel prototype was characterized by coil-current, electric-field, and thermal measurements. Results: The prototype produced a peak primary electric field of approximately 1 V/m measured in air 4 cm from the inner helmet surface. Whole-array modeling attained cortical fields of up to 1.5 V/m, reproduced the field distribution of a clinically validated low-intensity stimulator to within 3%-5%, and demonstrated focal targeting of the dorsolateral prefrontal cortex, simultaneous delivery of electric field to the default mode network nodes, and synthesis of electric fields following the traveling alpha wave. Conclusion: Electronically steerable, whole-head LI-TMS is feasible using accessible microprocessor-controlled power electronics. Significance: The array reaches the cortical field regime of tES without scalp contact or the associated shunting of current through the scalp, offering a route to testing network-level, phaselocked weak-field neuromodulation.
De Lillo, F.; Smucler, J.
Show abstract
Electrical stimulation (ES) and transepithelial/transendothelial electrical resistance (TEER) measurements are essential techniques in cell biology and tissue engineering, yet commercial devices for these applications cost between USD 2,500-9,000 and typically offer only one functionality. We present LATEER (Low-cost Arduino-based TEER and Electrical stimulation device), an open-source hardware platform that combines both ES and TEER measurement capabilities at a total cost below USD 100. The device features four independent channels, configurable pulsatile signals (amplitude up to 8.2 V, frequency 0.1-500 Hz, pulse width [≥]0.1 ms), and a resistance measurement range of 300 {Omega} to 1 M{Omega}, with <5% error for R {gtrsim} 4.7 k{Omega}. LATEER uses commercially available graphite pencil leads as electrodes ([~]USD 2 vs. USD 350 for commercial Ag/AgCl electrodes), which demonstrated excellent biocompatibility in cell culture. The system includes 3D-printed electrode holders compatible with standard 12-well and 24-well plates, allowing microscope visualization without electrode removal, and a Python-based graphical user interface for parameter configuration and real-time data acquisition. Because the electrodes remain fixed in the plate lid and only a single cable enters the incubator, both stimulation and resistance measurement can run continuously under standard culture conditions (37 {degrees}C, 5% CO2) without removing the plate or repositioning the electrodes, avoiding the temperature excursions and placement variability inherent to manual chopstick measurements. Validation with human pluripotent stem cell-derived cardiomyocytes demonstrated reliable frequency capture (electrical pacing) of the contracting monolayer, with a capture threshold between 250 and 400 mV/mm and controlled pacing across the 0.5-5 Hz range. TEER functionality was verified with mesenchymal stem cells, where the device resolved cell-density-dependent differences in electrical resistance in real time. All design files, firmware, and software are freely available under the CERN-OHL-S v2 license, enabling replication and customization by research laboratories worldwide. HighlightsO_LIAn open-source device combines electrical stimulation and TEER measurement under $100 C_LIO_LIGraphite electrodes offer biocompatibility at 0.6% cost of commercial alternatives C_LIO_LIFour independent channels with configurable parameters and real-time data logging. C_LIO_LIContinuous run setup in-incubator; no electrode repositioning needed C_LIO_LIValidated with stem cell-derived cardiomyocytes, achieving frequency capture (threshold 250-400 mV/mm) C_LIO_LI3D-printed holders enable microscope visualization without electrode removal C_LI Graphical abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=78 SRC="FIGDIR/small/743263v1_ufig1.gif" ALT="Figure 1"> View larger version (28K): org.highwire.dtl.DTLVardef@98f9deorg.highwire.dtl.DTLVardef@13c73aborg.highwire.dtl.DTLVardef@1cdf099org.highwire.dtl.DTLVardef@16ed0cf_HPS_FORMAT_FIGEXP M_FIG C_FIG Specifications Table O_TBL View this table: org.highwire.dtl.DTLVardef@4ef802org.highwire.dtl.DTLVardef@7c651borg.highwire.dtl.DTLVardef@d20013org.highwire.dtl.DTLVardef@102fc89org.highwire.dtl.DTLVardef@111b927_HPS_FORMAT_FIGEXP M_TBL C_TBL
Fujibuchi, T.
Show abstract
Reported relative biological effectiveness (RBE) values for low-energy X-rays disagree, assays scoring initial DNA double-strand breaks (DSBs) returning about 1.1 and chromosome-level assays 2 to 4. Whether radiation quality varies within the diagnostic range, and how its comparison with a megavoltage reference depends on target scale, has not been quantified on a tube-potential series. A tungsten-anode tube with 1 mm Be and 2.5 mm Al filtration, with copper added in some cases, was modelled in PHITS for 40 to 200 kV. The spectra were transported into a water phantom in which absorbed dose, lineal-energy densities and cluster size distributions were scored for target diameters of 3 nm to 1 micrometre against a cobalt-60 reference; DSB yields were computed in the electron track-structure mode with the PHITS DNA damage tally. Between 40 and 120 kV the depth-dose ratio changed by a factor of 5.7 and the tube output by a factor of 42, whereas the dose-mean lineal energy varied by 2.5 % at 1 micrometre and 1.2 % at 3 nm against a reproducibility of 0.3 %. Relative to cobalt-60 it was 2.05 times larger at 1 micrometre but only 1.08 times larger at 3 nm, while DSB yields per unit dose were 5 to 7 % higher and constant across the range within the 2 % bound set by the statistics. Tube potential therefore changes the amount and distribution of dose but not its physical quality, and a stated RBE is incomplete without the target scale implied by the endpoint.
Fleming, M. R.; Tayon, K. G.; Schneider, A.; McPherson, A. D.; Bianco, S. M.; Parent, E. E.; Sharma, A.; Lin, G.; Norton, N.; Ray, J. C.
Show abstract
Background. Cardiovascular disease is a leading cause of death among women with breast cancer, and the 2026 ACC/AHA dyslipidemia guideline endorses coronary artery calcium (CAC) scoring to guide statin therapy before cardiotoxic treatment. Breast cancer patients routinely undergo staging 18F-fluorodeoxyglucose PET/CT, whose low-dose CT visualizes the coronary arteries, thus enabling CAC quantification at no additional cost or radiation. Methods. In this single-center retrospective study, consecutive women with newly diagnosed breast cancer undergoing staging 18F-FDG PET/CT (2009?2021) had semi-automated Agatston CAC scoring performed on the low-dose CT and were stratified by CAC presence (CAC-P) versus absence (CAC-A). We assessed a composite of cardiac diagnostic testing (stress testing, coronary CT angiography, invasive angiography), clinical events, and reclassification of statin eligibility per ACC/AHA guideline thresholds in a prevention-eligible subgroup. Results. Among 276 women (mean age 55.5 years; median follow-up 7.1 years), CAC was present in 68 (25%) but was clinically reported in only 5.4%. CAC-P was associated with more cardiac testing (34% vs 12%; age-adjusted hazard ratio 2.75, 95% CI 1.43?5.28) and, though underpowered, with more atherosclerotic events (7.4% vs 1.4%), but not with the all-cause composite. In the prevention-eligible subgroup (n=39), CAC scoring would have changed statin eligibility in 64%, initiating therapy in 62% of CAC-P women and supporting de-prescribing in 67% of CAC-A women. Conclusions. CAC can be feasibly quantified from staging PET/CT in women with breast cancer and would frequently reclassify statin eligibility at no additional cost or radiation, yet is rarely reported.
Le Guellec, B.; Bentegeac, R.; Tran, V.-T.; El Homsi, M.; Amouyel, P.; Kuchcinski, G.; Hamroun, A.
Show abstract
Background: Large language models have been proposed to improve patient comprehension of radiology reports. However, whether they improve objective understanding remains unproven. Purpose: To evaluate the effect of appending an LLM-generated lay summary to brain MRI reports on objective and subjective patient comprehension in a randomized controlled trial. Materials and Methods: In this randomized controlled trial, 2,727 adult participants from the ComPaRe e-cohort were randomly assigned to interpret six standardized brain MRI reports for headache, presented either in their native format (control; n = 1,401) or appended with a lay summary generated by an open-weights LLM (Mistral Small 3.2) (intervention; n = 1,326). The primary outcome was objective comprehension, defined as the rate of correct classification of whether the report provided a probable explanation for the headache, with ground truth established by four-radiologist consensus. Secondary outcomes included satisfaction, subjective comprehension, anxiety, and willingness to contact a healthcare professional. Generalized estimating equations accounted for repeated within-participant observations. Results: A total of 2,727 participants (mean age, 52 years +/- 15; 75.2% women) were evaluated. Objective comprehension did not differ between arms (58.3% vs 59.4%; odds ratio (OR) 0.97; 95% CI: 0.90-1.06; P = .54). The intervention significantly improved overall satisfaction (64.9% vs 36.7%; OR 3.26; 95% CI: 2.93-3.64; P < .001) and subjective comprehension (50.3% vs 24.0%; OR 3.17; 95% CI: 2.82-3.56; P < .001). High anxiety was modestly reduced (25.1% vs 26.6%; OR 0.92; P = .037). The effect on objective comprehension varied by report type (P for interaction < .001): summaries improved comprehension of symptom-explaining reports (42.4% vs 37.4%; P < .001) but reduced it for normal reports (72.5% vs 76.6%; P = .001). Conclusion: LLM-generated lay summaries appended to brain MRI reports improved patient satisfaction and subjective comprehension but did not improve objective comprehension, indicating a gap between perceived and actual understanding that should be addressed before clinical integration.
Ivanov, B.; Arvaneh, M.; Toth, J.; Rampersad, S. M.
Show abstract
AbstractComputational models of temporal interference stimulation (TIS) commonly report a single electric-field estimate for a given anatomy and electrode montage. Because non-deterministic tetrahedral mesh generation does not produce a unique discretisation of a fixed tissue-label image, a single mesh realisation may introduce numerical variability. We quantified variation across independent mesh realisations and contrasted it with repeated downstream simulation execution on a single selected mesh. Ten head models were evaluated for stimulation of the left hippocampus and right primary motor cortex (M1). For every model and target, we generated 40 independent meshes and performed one complete simulation on each. Separately, we selected the mesh whose parcel-level field estimate was closest to the median and repeated downstream operations 40 times while holding that geometry fixed, yielding 1,600 TIS simulations in total. The primary outcome was the spatial median of the TIS envelope field within a spherical target region. Across independently remeshed runs, within-participant coefficients of variation were 1.81-3.65% for the hippocampus and 1.62-2.79% for M1. Repeated execution on a fixed mesh reduced run-to-run standard deviation by more than 99%, demonstrating that workflow variability is driven almost entirely by non-deterministic mesh generation rather than solver instability, numerical rounding, or post-processing. Single-run mesh realisations preserved overall cohort ordering (median Kendalls{tau} of 0.867 for the hippocampus and 0.911 for M1) but frequently inverted the rank order of participant pairs with similar predicted fields. Furthermore, a bootstrap analysis demonstrated that averaging five to ten independent remesh runs effectively suppressed this stochastic noise. These results quantify single-workflow repeatability rather than absolute error. Stochastic mesh variation should therefore be controlled or mitigated through multi-run averaging whenever experimental conclusions depend on subtle field differences or fixed neuromodulation thresholds.
Lapatrie, M.; Isetani, Y.; Puvirajan, J.; Catanzaro, A.; Lyu, S.; Nguyen, H. C.; Mathieu, W.; Popovic, M.
Show abstract
Transcranial magnetic stimulation (TMS) excites neurons noninvasively by electromagnetic induction and is used in neurophysiology research and in approved therapy for depression. Commercial stimulators cost tens of thousands of dollars. Existing open-source designs are either low-energy and unvalidated or rely on expensive switches and laboratory infrastructure. We present a monophasic, fixed-pulse-shape TMS device built at a parts cost of ~USD 700 which, under specific modeling assumptions, can exceed average human motor thresholds. Our design assumes access to basic, off-the-shelf equipment such as a 24 V power supply unit, an oscilloscope, and a few basic tools. The device charges a 230 F film-capacitor bank and discharges it through a self-wound figure-of-eight coil using a thyristor, producing a fixed pulse with a positive lobe lasting approximately 90 s. A Zero-Voltage Switching (ZVS) driver-based charging circuit charges the capacitor bank up to 1460 V from a 24 V bench supply. Three galvanically isolated voltage domains, redundant interlocks, and passive and active discharge paths help mitigate the safety risks involved with handling lethal energy levels. We also present a low-cost way to characterize the device by reconstructing coil di/dt from pickup-coil dB/dt maps to estimate the induced cortical E-fields. At the maximum capacitor voltage, the recovered maximal di/dt is 110.86 A/s, giving estimated 99.9th percentile cortical E-fields of 159 V/m at Oz and 196 V/m at C3 on an example anatomy. Although not yet approved for clinical trials and routine stimulation, the device demonstrated the possibility of a cost-effective TMS unit.
Chau, G. N.; Biswas, B. A.; Wagle, B. R.; Maeder, M. E.; Yu, J. B.; Bhattacharya, I.
Show abstract
Automated lesion segmentation is increasingly central to PSMA PET/CT interpretation, supporting staging, treatment planning, and response assessment at a scale that outpaces available nuclear-medicine expertise. However, automated PSMA-PET/CT whole-body lesion segmentation models are trained on images alone, with no knowledge of where in the body prostate metastases actually tend to occur. Radiologists use clinical domain knowledge of metastatic spread, but its absence in machine learning models produces false positives in anatomically implausible locations and missed lesions in high-risk sites such as the liver. In this work, we explore whether population-level spatial knowledge of metastatic spread can be used to augment deep learning segmentation predictions, and how such a prior should be fused with a network's output, without additional training. We build a data-driven metastasis atlas from 375 expert-annotated whole-body PSMA PET/CT scans and investigate its fusion with a trained segmentation network under a Bayesian framework, in which prediction probabilities from an nnU-Net-based lesion segmentation model serve as the likelihood and the data-driven atlas as the prior. Because metastases occupy only a small fraction of whole-body voxels, the atlas's peak probability is too low, and standard power-scaled or naive Bayesian pooling references lack the tools to deal with this shortcoming. This causes these standard fusion strategies to fail and, in the naive Bayesian case, to sharply degrade performance. We instead derive a calibrated, background-referenced log-odds fusion, one of many possible approaches to combine a population atlas with a deep learning model's predictions, distinct from classical multi-atlas label fusion in that it fuses a single population prior with a trained network's softmax rather than combining several registered atlases. Furthermore, this approach is neutral outside atlas support by construction, reduces exactly to the baseline network when unweighted, and requires no retraining. This atlas fusion significantly improved mean Dice over the baseline nnU-Net on a disjoint internal test set ($+0.011$, Holm-adjusted $p=0.021$) and on an independent external cohort ($+0.0129$, Holm-adjusted $p=3.8\times10^{-16}$), with lesion sensitivity improving from 0.849 to 0.861 internally and Dice improving over baseline in every stratified anatomic region, including the rare, high-risk sites motivating this work, while naive Bayesian pooling degrades performance sharply and power-scaled pooling underperforms it throughout. Our findings suggest that population-level spatial priors can meaningfully augment deep learning predictions in whole-body oncologic segmentation, provided the fusion rule is calibrated to where the prior actually carries signal.
BAI, T.-C.; YEH, S.-C.
Show abstract
CXR report generation may require a vision-language model (VLM) to produce both textual findings and spatial bounding boxes. Generative 4B-7B VLMs can emit non-empty outputs on normal images and empty outputs on abnormal images, motivating explicit structural routing. To evaluate whether a hard inference-time gate before a probabilistic VLM changes output-presence performance and to identify the mechanisms underlying paired STRUCT outcomes. We evaluated CXRxVLM v2, combining a frozen microsoft/rad-dino ViT-B/14 encoder with a 768[->]1 logistic probe (threshold 0.0557) and google/medgemma-4b-it with the pamessina/medgemma-4b-it-cure LoRA adapter. A seed=42 stratified cohort of 500 VinDr-CXR train-pool images (250 NORMAL, 250 ABNORMAL) was compared with Lingshu-7B A_baseline and D_fewshot configurations. Exact paired McNemar tests and stratum-level output-presence analyses were prespecified for the primary configurations; MedGemma 1.5 SigLIP was exploratory. CURE achieved STRUCT = 78.0% (390/500; Wilson 95% CI 74.2-81.4), versus 73.8% for Lingshu A_baseline and 74.2% for D_fewshot. Pairwise p-values were 0.0778, 0.1042, and 0.8642. The paired decomposition showed CURE ABNORMAL non-empty-output advantage of +13.6 percentage points versus Lingshu A (p = 0.0012; +14.0 points versus D, p = 0.0007), while Lingshu had higher NORMAL empty-output rates (+5.2 to +6.4 points; p = 0.0106 and p = 0.0004). The full pipeline used 8.87 GB VRAM and 4.92 s/image mean latency; 53% of records used a 25.7 ms warm gate-negative path after model loading. Equivalent overall STRUCT scores concealed two mechanistically different output regimes: CURE favored ABNORMAL non-empty outputs, whereas Lingshu favored NORMAL empty outputs. This paired decomposition, rather than the aggregate score alone, characterizes how hard-gated and probabilistic systems route output presence.
Khadka, N.; Huang, Y.; Deng, Z.-D.; Truong, D. Q.; Venkatasubramanian, G.; Tu, Y.; Ma, W.; Abbott, C. C.; Datta, A.
Show abstract
Objective: This computational modeling study quantified the influence of sex and race-related cranial anatomy on predicted brain-wide current flow during electroconvulsive therapy (ECT) across conventional (bifrontal (BF), bitemporal/bilateral (BL), right unilateral (RUL)) and experimental (focal electrically administered seizure therapy (FEAST) and frontomedial (FM)) electrode montages. The objective was to determine whether race-associated variability meaningfully contributes to differences in ECT stimulation metrics across montages. Methods: Finite element head models of Chinese, Black, and Caucasian subjects were developed using high-resolution magnetic resonance imaging and analyzed using the Realistic vOlumetric- Approach-based Stimulator for Transcranial electric stimulation (ROAST) pipeline (N = 150 total; n = 50 per cohort, comprising 25 M and 25 F, age range: 20-30 years). Five ECT montages were simulated under a constant-current condition (900mA). Stimulation strength (Ebrain/Eth) was quantified as 90th percentile of brain-wide E-field magnitude (Ebrain) relative to neuronal activation threshold (Eth = 0.25 V/cm) quantified stimulation strength. Overall focality was evaluated as a percentage of brain volume stimulated above the neural activation threshold (Ebrain [≥] Eth), while laterality was quantified as the median right-to-left hemispheric E-field magnitude ratio. The effects of race, sex, and montage on stimulation strength, focality, and hemispheric laterality were statistically analyzed. Results: Substantial race- and sex-related differences observed in cranial anatomy resulted in systematic variation in predicted ECT-induced E-field intensity. Brain-wide E-field magnitude varied by both race and montage, with the largest fields generally observed in Caucasian head models and during BL stimulation. Montage exerted the strongest effect on stimulation strength (Ebrain/Eth) with BL and FEAST producing the highest stimulation strengths, followed by RUL and FM, while BF produced the lowest. Caucasian subjects generally predicted higher stimulation strengths than Black and Chinese subjects, whereas females predicted modestly higher stimulation strengths than males. Laterality was primarily determined by montage, with FEAST producing the greatest hemispheric asymmetry, followed by RUL. Chinese subjects demonstrated higher laterality ratios than both Black and Caucasian subjects. BL, RUL, and FEAST stimulated substantially larger brain volumes above neural activation threshold (less focal stimulation) than BF. Lower focality was observed in Caucasian subjects relative to Black and Chinese subjects, and in females relative to males. Conclusions: Electrode montage was the primary determinant of predicted ECT stimulation strength, focality, and laterality. Race-related anatomical differences and, to a lesser extent, sex-related differences systematically altered stimulation patterns, supporting consideration of individualized anatomy in ECT dosing and treatment optimization.
Courtens, J.; Muller, F. M.; Li, E. J.; Vanhove, C.; Vandenberghe, S.; Pantel, A. R.; Karp, J. S.; Daube-Witherspoon, M. E.
Show abstract
Dynamic positron emission tomography (PET) with long axial field-of-view (LAFOV) scanners enables multi-organ imaging and kinetic quantification beyond static (late-phase) imaging; however, the long times typically required for dynamic acquisitions remain clinically impractical. This study evaluates a deep learning (DL) framework to enable abbreviated dynamic PET acquisitions, comparing single-time-window (STW, early dynamic data only) and dual-time-window (DTW, early dynamic data plus a late 5-min static frame) protocols with early dynamic scan durations of 5-30 min and dose levels ranging from 360 MBq to 18 MBq. Seventeen 60-min dynamic [18F]FDG datasets were first motion-corrected using a staggered FALCON pipeline and then used to train and test a spatiotemporal DL model for autoregressive frame prediction. Performance was assessed across the full quantitative workflow, from DL-predicted frames and time-activity curves to organ-based kinetic modeling and voxel-wise parametric imaging in multiple tissues and two patient cohorts. DTW protocols consistently outperformed STW, better preserving late-phase kinetics. For a 15-min early dynamic scan, adding a late 5-min scan reduced mean absolute Ki difference from 23% (STW) to 17% (DTW) in the liver and from 26% to 15% in the thalamus. DTW + DL further reduced errors to [≤]10% in the liver, thalamus, and breast lesion, and 16% in muscle. Our recommended protocol, 15-min early dynamic scan plus a 5-min late scan with DL, remained robust to up to a 5-fold dose reduction (~74 MBq). Overall, these findings support DL-enabled abbreviated, low-dose dynamic LAFOV PET as a clinically feasible approach for accurate kinetic quantification
Bartels, R.; Vinke, S.; Rijpma, A.; Nadimi, M.
Show abstract
Deep brain stimulation (DBS) modeling relies heavily on biophysical neuron models to estimate neural activation thresholds and predict stimulation spread. In this study, we systematically compared a widely adopted axon model, the McIntyre-Richardson-Grill (MRG) model (Model I), with a more detailed biophysical model, the Cohen model (Model II), to assess how structural and electrophysiological differences affect predicted DBS outcomes. Electric field distributions generated by 2202 DBS lead were applied to the neuron models as extracellular input stimuli. Both models were simulated under biphasic pulse stimulation across varying axon-electrode distances, pulse widths, and stimulation frequencies. Activation distances ranged from approximately 2 to 10 mm depending on stimulation parameters and contact location. At 2 mA, Model I achieved an activation distance of 6 mm, whereas Model II reached 10 mm, indicating greater excitability. Across matched fiber tracts, threshold differences ranged from -1.40 mA to 0.27 mA, with Model II requiring lower thresholds in 97.7% of cases. Both models showed a strong inverse relationship between pulse width and activation threshold. However, frequency responses differed: Model II exhibited increasing thresholds at higher frequencies, while Model I showed a slight decrease. Machine learning regressors trained on distance, pulse width, and frequency achieved high predictive accuracy, with Gradient Boosting performing best. Model II demonstrated superior prediction metrics (R^2 = 0.986; RMSE = 0.045 mA; MAE = 0.034 mA) compared to Model I (R^2 = 0.977; RMSE = 0.089 mA; MAE = 0.068 mA). Overall, both models reliably estimate DBS-induced activation, but structural differences significantly affect excitability and frequency-dependent behavior. With appropriate awareness of their respective strengths and limitations, either model can be used to derive activation distances for estimating electric field isolevels and the volume of tissue activated in patient-specific DBS simulations.
Bagchi, R.; Yee, N. J.; Kwon, J. Y.; Taseh, A.; Ashkani-Esfahani, S.
Show abstract
Purpose To evaluate whether domain-adaptive self-supervised pretraining on musculoskeletal radiographs improves fracture classification and attribution faithfulness relative to ImageNet-pretrained baselines. Materials and Methods This study (June 2025 to May 2026) used previously acquired radiographs to compare three ResNet-50 initializations: supervised ImageNet pretraining (control), self-supervised ImageNet pretraining (DINO), and DINO with additional domain-adapted pretraining on 44,029 musculoskeletal radiographs (DINO-Ortho). All models underwent supervised fine-tuning in three experiments: in-distribution (MURA and FracAtlas datasets), out-of-distribution (an external dataset of 5,365 calcaneal radiographs from 1,775 patients), and initial weights (calcaneal radiographs only). Metrics included sensitivity, specificity, test accuracy, area under the receiver operating characteristic curve (AUROC), and Cohen's kappa; attribution faithfulness was quantified using Remove and Debias scores from Grad-CAM saliency maps. Comparisons used DeLong and Friedman tests. Results Classification performance did not differ significantly between DINO-Ortho and either baseline in any experiment (DINO-Ortho AUROC, 0.89 in-distribution and 0.95 with initial weights). All three models discriminated poorly out-of-distribution (control, 0.59; DINO, 0.57; DINO-Ortho, 0.58). DINO-Ortho showed significantly higher attribution faithfulness than both baselines in all three experiments, including out-of-distribution (25.39 vs -10.41 and 2.14; P < .001) and initial weights (20.88 vs 11.51 and 1.27; P < .001). Qualitative rankings favored DINO-Ortho but did not differ significantly. Conclusion Domain-adapted self-supervised pretraining on musculoskeletal radiographs improved attribution faithfulness while maintaining classification performance comparable to ImageNet-pretrained baselines; no model generalized adequately to external radiographs without task-specific fine-tuning.
Kenanoglu, C. U.; Vardar, Y.
Show abstract
Electrostatic actuation is an emerging technology for generating tactile sensations on capacitive touchscreens through voltage-induced attractive forces between a fingertip and the surface. However, accurate control of electrostatic attraction during natural touchscreen interactions remains challenging because the applied normal force and sliding speed continuously vary, and their effects on the fingertip-screen contact and resulting actuation strength are not fully characterized. Here, we show how normal force and sliding speed systematically alter fingertip- screen contact area and electrical impedance, and use these measured changes to estimate electrostatic attraction during sliding. Contact area, interaction forces, and electrical impedance were measured simultaneously as participants slid their fingertips across an electrostatic surface under systematically varied normal forces and sliding speeds. These measurements revealed condition-dependent changes in fingertip contact, electrical interaction impedance, effective capacitance, derived effective gap thickness, and electrostatic attraction. We then incorporated these measured contact quantities into a physics-informed, data-driven model based on parallel-plate capacitor theory, in which effective capacitance, apparent contact area, and effective voltage determine the estimated electrostatic attraction. The resulting model links force- and speed-dependent changes in these quantities to electrostatic attraction while accounting for inter-participant variability through a participant-specific scaling factor. These findings provide experimentally grounded guidance for designing electrostatic surface-haptic feedback and future adaptive control strategies under realistic touch conditions.
Dack, E.; Dai, C.; Hoppe, H.; Krueselmann, P.; Meiler, S.; Jutidamrongphan, W.; Wang, L.; Tang, K.
Show abstract
AI-assisted diagnostic tools typically act as a "second opinion," providing radiologists with a discrete prediction or probability score that can be consulted alongside clinical context. This treats AI as an independent advisor rather than a collaborative partner, leaving its reasoning largely opaque. We explore a complementary approach grounded in human-AI collaboration through visual interpretability. Specifically, we investigate (1) radiologist performance when diagnosing chest X-rays from images alone, and (2) whether deep learning-generated heatmaps can support radiologists during this diagnostic process, rather than merely validating a final answer. We developed an interactive application that enables readers to engage directly with model-generated heatmaps as they form their diagnoses, and conducted a user study to evaluate how this influences diagnostic behaviour and accuracy. Our findings offer new insights into integrating interpretable, spatially grounded AI feedback into radiologist workflows. Code, datasets, and the application can be found at https://github.com/eedack01/heatmap_assisted_diagnosis.